Viewpoint
Abstract
Cancer diagnosis depends on data from radiology, digital pathology, molecular profiling, laboratory testing, and longitudinal clinical records. AI performs well in selected tasks, but most systems remain narrow and disconnected from the iterative reasoning required in oncology. This Viewpoint defines an AI agent as a feedback-driven system that maintains task state, selects among governed tools, observes results, and revises its plan under explicit safety constraints. This definition separates agents from multimodal foundation models, retrieval-augmented generation, and fixed workflow automation. We organize the discussion across multimodal data collection, preprocessing, fusion and representation learning, and diagnostic decision support. We distinguished agent-level evidence, component- or infrastructure-level evidence, and prospective propositions throughout. Clinical translation will require resilient failure handling, guideline version control, prospective evaluation, computational and workflow feasibility, and clinician authority over final decisions. The near-term opportunity is therefore transparent and traceable clinical decision support rather than autonomous cancer diagnosis.
JMIR Cancer 2026;12:e103545doi:10.2196/103545
Keywords
Introduction
Cancer remains a leading cause of morbidity and mortality. Global Cancer Observatory (GLOBOCAN) 2022 estimated that there were nearly 20 million new cancer cases and 9.7 million cancer-related deaths worldwide in that year []. Treatment delays are associated with higher mortality across several cancer indications []. Timely diagnostic pathways are therefore important, but they must also account for genomic, histopathological, microenvironmental, and temporal heterogeneity [].
Modern oncology diagnosis is intrinsically multimodal. Clinicians combine imaging, pathology, laboratory results, molecular assays, and longitudinal electronic records, often across separate specialties and information systems. This fragmentation creates cognitive and coordination burdens that motivate integrated workflows [,]. Task-specific AI has advanced lesion detection, pathology classification, molecular modeling, and risk prediction [,-]. However, most deployed or evaluated models address a single bounded input-output task rather than the full diagnostic process.
Large language models (LLMs), multimodal foundation models, and tool-use frameworks create a technical basis for coordinating these tasks [,]. However, a model that accepts several modalities is not automatically an agent. Similarly, retrieval-augmented generation (RAG) adds external evidence to generation, and fixed workflow orchestration links predetermined modules, but neither property alone establishes runtime planning or feedback-driven control.
This Viewpoint examines AI agents as auditable orchestration layers for multimodal oncology diagnostics. We organized the field around 4 connected stages: data collection, preprocessing, multimodal fusion and representation learning, and diagnostic decision support. Our central position is deliberately bounded: agentic orchestration may improve coordination and auditability, but clinical benefit and safety remain deployment-specific questions.
From Task-Specific AI to Agentic Oncology Diagnostics
We used an operational definition of agentic behavior. In a conventional fixed pipeline, an output is generated through a predetermined composition, y=fn(...f₁(x)), and the sequence of functions does not change at runtime. An AI agent instead maintains a state st and selects an action at = π(g, st, At) from an allowed tool set At. For a clinical goal g, it observes each tool’s result and updates its state before choosing the next action. The practical distinction therefore lies in conditional planning, tool selection, observation, and revision within a governed loop.
Four adjacent concepts should remain separate. A multimodal foundation model supplies reusable representations or generation across data types. RAG retrieves external sources to ground an answer. Fixed workflow orchestration executes a prescribed sequence of services. A feedback-driven AI agent can change the sequence, request missing evidence, retry or switch tools, abstain, and escalate to a clinician according to the evolving state. Long-term memory is optional and should be limited to an authorized, versioned, patient-specific or institutional state rather than treated as an unrestricted capability.
First, contrasts a conventional specialist synthesis with an agent-coordinated 4-stage workflow. Second, operationalizes the agentic layer as a plan-act-observe-update-check loop. Act coordinates the 4 stages in B; update records state and provenance; and check evaluates evidence sufficiency, cross-modal consistency, and whether further action remains within the authorized scope. If these criteria are met, the synthesis proceeds to clinician review. Remediable gaps return the workflow to planning; gaps that cannot be supplemented trigger clinician escalation. The agent-coordinated elements in both figures are prospective design propositions and not validated clinical systems or performance claims.
We applied 3 evidence labels throughout this Viewpoint. Agent-level evidence evaluates a complete loop that plans, uses tools, and produces an integrated output, as in the external oncology evaluation by Ferber et al []. Component or infrastructure evidence evaluates a model, retrieval method, data pipeline, or governance mechanism that an agent could invoke []. Prospective propositions describe plausible clinical workflows that still require system-level and prospective validation. This hierarchy prevents strong component performance from being misread as proof of a clinically validated agent.


Multimodal Data Collection
The first stage is the governed assembly of patient-specific evidence. Relevant data may be distributed across a hospital information system (HIS), laboratory information system (LIS), picture archiving and communication system (PACS), electronic medical record (EMR), pathology archives, and genomic platforms. These systems differ in format, ontology, temporal granularity, and access control.
Component-level infrastructure studies show that integration must be engineered. A single-center oncology data supply chain linked clinical, genomic, and imaging data for more than 170,000 patients across 11 cancer types and more than 800 features per case []. Its extract-transform-load (ETL), natural language processing (NLP), and quality control procedures support the feasibility of a data layer but not the validation of an agent. Standards such as Health Level Seven Fast Healthcare Interoperability Resources (HL7 FHIR) and Digital Imaging and Communications in Medicine (DICOM) can reduce site-specific translation []. Local mapping and governance nevertheless remain necessary.
An agent could use secure interfaces to identify, retrieve, and temporally align data for a specific diagnostic question. It could connect a new scan to prior pathology, biomarker trajectories, sequencing results, and notes [,]. Unstructured text and patient-reported or wearable data may add context in selected settings []. Each input should carry its source, acquisition and update times, missingness status, and freshness threshold. Contradictory or stale records should trigger clarification rather than silent fusion. Once evidence is assembled, the next risk is whether every modality is technically fit for analysis.
Multimodal Data Preprocessing
Multimodal preprocessing converts heterogeneous inputs into representations that downstream tools can use. Imaging varies by scanner and acquisition protocol. Digital pathology depends on tissue preparation, staining, and slide digitization []. Laboratory and genomic data may be incomplete or generated on different platforms, while clinical text often contains duplication and institution-specific shorthand.
A fixed workflow applies the same preprocessing sequence. An agentic workflow could select modality-specific tools according to file type, quality, and clinical purpose. It might invoke imaging quality control and segmentation or pathology stain normalization, tiling, tissue filtering, and feature extraction. Pathology foundation models provide reusable component representations [,], while self-supervised medical imaging models illustrate robustness and data-efficiency strategies []. Attribution tools such as gradient-weighted class activation mapping can support region-level inspection, although an attribution map is not itself a clinical explanation [].
Every action should be recorded with the input source, time stamp, parser or model version, quality control result, and review status. If optical character recognition or parsing fails, the system should not fabricate a normalized record. It should retry within a bounded policy, switch to a validated alternate parser, return a partial result with an explicit missingness flag, or request human review. This provenance-aware preprocessing creates the conditions for fusion without concealing upstream uncertainty.
Multimodal Fusion and Representation Learning
Fusion is not the mere presence of multiple data types. Imaging describes macroscopic phenotype, pathology reveals tissue architecture, molecular assays characterize biological drivers, and longitudinal records provide clinical context. Multimodal models can combine these views for characterization and risk estimation [], but many assume that all modalities are harmonized and simultaneously available.
At the molecular level, machine learning can integrate complementary omics layers and prioritize regulatory interactions. Omidi combined RNA sequencing and microRNA sequencing profiles from 2 public cohorts []. The analysis included 78 adrenocortical carcinoma samples from The Cancer Genome Atlas adrenocortical carcinoma (TCGA-ACC) and 250 normal adrenal samples from the Genotype-Tissue Expression (GTEx) project, reconstructing condition-specific competing endogenous RNA networks. Random forest models performed best for predicting microRNA-messenger RNA associations, and predicted interactions were cross-referenced with TargetScan and miRTarBase. This study is an in silico systems biology module that an agent could invoke. Batch effects, cohort imbalance, cross-cohort heterogeneity, and the absence of independent or experimental validation prevent these findings from being interpreted as validation of a multimodal agent.
An agent could select a fusion strategy according to the clinical question, available modalities, and uncertainty. If pathology and imaging are available while genomic testing is pending, it could use a modality-robust model, mark molecular evidence as unavailable, and revise the synthesis later. Explainable multimodal real-world models have identified patient-level marker contributions in large cohorts and undergone independent validation []. These results support component-level representation and explanation but not autonomous clinical orchestration. Fusion outputs must therefore pass an evidence-adequacy check before decision support. The evidence-adequacy gate and its return-or-escalate logic are shown in .
Diagnostic Decision Support
The final stage translates analyzed evidence into a structured, reviewable diagnostic summary. The purpose is not autonomous diagnosis or treatment selection. It is to state the supported findings, unresolved conflicts, missing evidence, uncertainty, and provenance in a form that clinicians can audit.
A recent feasibility study evaluated Gemini 2.5 Pro (Alphabet Inc) and ChatGPT 4o (OpenAI) on 80 Turkish-language prostate-specific membrane antigen positron emission tomography–computed tomography (PSMA PET-CT) reports []. The prompts incorporated the American Joint Committee on Cancer (AJCC) staging, the Chemohormonal Therapy Versus Androgen Ablation Randomized Trial for Extensive Disease (CHAARTED) criteria, and few-shot examples. The models assigned T, N, and M categories and disease volume classes, with overall task accuracies of 93.8% and 91.3%, respectively. Errors clustered around equivocal findings and different interpretive thresholds, including overstaging of ambiguous lesions. The single-center design, 1 expert reference standard, and exploratory sample of 80 reports limit generalizability. The study supports an LLM-based report-to-staging component that requires expert oversight; it does not validate a complete AI agent.
Other component studies illustrate evaluable end points. Multimodal oncology chatbots have been tested on image-containing clinical cases []. Explanation types can alter physician trust and diagnostic performance []. On the contrary, fabricated details can induce hallucinated clinical elaboration despite simple mitigation attempts []. Agent evaluation should therefore include tool selection, evidence citation, contradiction handling, safe refusal, calibration, workflow burden, and final conclusion quality rather than fluency alone.
In practice, an agent might prepare a multidisciplinary team (MDT) case by linking a suspicious lung lesion to pathology, molecular markers, smoking history, and prior imaging. It could rank competing diagnoses, identify a missing confirmatory test, and provide an evidence-linked summary. Proposed MDT assistance and randomized source-attribution studies support case preparation and verifiable retrieval as useful components [,]. Clinicians must retain authority over the final judgment, with explicit responsibility assigned across users, institutions, and manufacturers [].
Challenges and Future Directions
Limits of the Current Evidence
The evidence base remains uneven. Most cited studies validate models, infrastructure, or simulated workflows rather than prospectively deployed agentic systems. The following implementation requirements should therefore be read as a translational agenda and not as established clinical effectiveness.
Data Quality and Interoperability
Interoperability does not guarantee semantic correctness. A deployable system must map local fields to versioned ontologies, record lineage and freshness, and quantify missingness. It must also test performance across scanners, laboratories, institutions, demographic groups, and unavailable-modality patterns [,,]. External validation and subgroup analysis remain necessary because an agent can amplify errors introduced by any upstream component.
Operational Resilience and Graceful Degradation
Real-time MDT use requires an explicit latency budget for retrieval, preprocessing, inference, and rendering. Institutions may precompute stable features, use asynchronous tasks, or choose smaller local models when the expected benefit does not justify the delay. Time-outs should activate circuit breakers rather than indefinite retries. A parser or API failure should trigger a typed error and one of the following governed responses: a bounded retry, use of a validated alternate tool, a partial-result flag, or abstention with human escalation. The audit record should preserve the failed attempt and the fallback path. This fallback logic corresponds to the supplementation and clinician-escalation branches in .
Knowledge Life Cycle and Guideline Synchronization
Oncology guidelines and drug knowledge change frequently. Retrieval stores should preserve source dates, jurisdiction, guideline version, and effective period. Updates require staged ingestion, regression tests against reference cases, comparison with the previous version, canary deployment, and rollback criteria. A response should identify the version used and warn when local policy conflicts with an external guideline. This reduces, but cannot eliminate, legacy hallucinations.
Evaluation and Clinical Integration
Evaluation should follow the clinical workflow and the evidence label. Component metrics include extraction accuracy, calibration, source verifiability, latency, throughput, and failure rate. Agent-level evaluation additionally requires plan appropriateness, tool-use accuracy, recovery from failed steps, contradiction resolution, abstention, clinician workload, and patient safety end points [,-]. Prospective studies should compare the full human-agent workflow with current practice and not only the agent with a static model.
Governance, Cost, and Institutional Adoption
Deployment also depends on computing cost, local infrastructure, cybersecurity, procurement, workflow fit, regulatory status, reimbursement, and institutional adoption. Trustworthy AI principles require fairness, traceability, usability, robustness, and explainability []. Precision oncology governance should add access control, deidentification, audit logging, incident response, continuous monitoring, model and knowledge base change control, and explicit accountability [,]. These controls are prerequisites for a clinical service and not evidence that the service improves outcomes.
summarizes the current evidence supporting AI agents in oncology diagnostics and the limitations that constrain their interpretation and clinical use.
| Evidence domains | Study design or setting | Main empirical finding | Agentic relevance and boundary |
| Agent-level oncology decision support | External evaluation on 20 realistic multimodal oncology cases [] | Reported 87.5% correct tool use, 91% correct clinical conclusions, and 75.5% accurate guideline citation | Agent-level evidence from simulated cases; a useful benchmark but not prospective clinical validation |
| Real-world multimodal infrastructure | Single-center data supply chain covering >170,000 patients, 11 cancer types, and >800 features per case [] | Integrated clinical, genomic, and imaging data with extract-transform-load, natural language processing, quality control, and continuously updated records | Infrastructure evidence; supports a callable data layer but does not evaluate agent planning or clinical benefit |
| Multimodal oncology chatbot benchmarking | Case-based evaluation of multimodal chatbots on 79 image-containing oncology cases [] | Demonstrated that image-text oncology reasoning can be measured with case-level accuracy | Component evidence; a chatbot benchmark does not establish workflow orchestration or safety |
| Large language model–assisted prostate cancer staging from unstructured imaging reports | Single-center feasibility study of 80 Turkish-language prostate-specific membrane antigen positron emission tomography–computed tomography reports with embedded American Joint Committee on Cancer and Chemohormonal Therapy Versus Androgen Ablation Randomized Trial for Extensive Disease criteria and few-shot prompting [] | Overall task accuracy was 93.8% for Gemini 2.5 Pro and 91.3% for ChatGPT 4o; errors clustered around equivocal findings | Component evidence; supports a report-to-staging module requiring expert oversight and not a complete agent |
| Machine learning integration of molecular regulatory networks | In silico integration of 78 The Cancer Genome Atlas adrenocortical carcinoma tumors and 250 normal adrenal samples from the Genotype-Tissue Expression project using RNA and microRNA sequencing [] | Random forest best predicted microRNA-messenger RNA associations; predictions were cross-referenced with TargetScan and miRTarBase. | Specialized analytic module; cohort imbalance, batch effects, and the absence of independent or experimental validation limit inference |
| Tumor MDTa workflow | Workflow-focused proposal for AI assistance in tumor MDTs [] | Identified opportunities for case preparation, synthesis of unstructured data, and rationale capture | Prospective workflow proposition; final decisions remain clinician led and require comparative evaluation |
| Source-attributed uro-oncology decision support | Randomized reader evaluation of a retrieval-augmented uro-oncology chatbot [] | Improved recommendation correctness, source attribution, and source verifiability compared with a general chatbot | Component evidence; supports governed retrieval and verifiable sources and not full agent autonomy |
| Explainable multimodal real-world oncology AI | Multimodal model in 15,726 patients across 38 solid cancer types with external validation in 3288 patients with lung cancer [] | Estimated patient-level marker contributions and prognostic interactions | Model-level evidence; supports reusable prediction and explanation tools with external validity boundaries |
| Hallucination and safety | Multimodel simulation of adversarial hallucination during clinical decision support [] | Fabricated details induced false elaboration despite simple mitigation prompts | Safety evidence supporting contradiction checks, retrieval constraints, abstention, and human escalation |
| Interoperability | Clinical interoperability analysis across the AI data life cycle [] | Emphasized standardized information exchange for data entry, processing, sharing, replication, and analytics | Infrastructure evidence; local semantic mapping and provenance remain necessary |
| Governance, liability, and implementation | Consensus trustworthy AI guidance, liability analysis, and precision oncology implementation review [,,] | Identified requirements for fairness, traceability, usability, robustness, explainability, responsibility, safety, governance, cost, and harmonization | Cross-cutting requirements; governance controls enable evaluation but are not evidence of clinical effectiveness |
aMDT: multidisciplinary team.
Conclusions
AI agents should be understood as auditable, feedback-driven orchestration systems rather than as another name for multimodal models or RAG. Oncology-specific studies now support selected agent-level, component-level, and infrastructure capabilities, including tool use, report-to-staging, multimodal modeling, and source attribution. However, these findings do not establish the safety or benefit of autonomous diagnosis. The defensible near-term goal is a resilient human-agent workflow that exposes provenance, uncertainty, failures, guideline versions, and escalation decisions. Prospective comparison with current clinical practice is required before broader implementation claims are justified.
Acknowledgments
The authors thank colleagues and collaborators for valuable discussions and conceptual input and acknowledge institutional research platforms that facilitated literature retrieval and manuscript preparation. No other individuals or organizations contributed in a manner that met authorship criteria.
ChatGPT (version 5.6; OpenAI) was used during revision to support language polishing. The authors reviewed and approved all AI-assisted revisions and take full responsibility for the manuscript.
Funding
This work was supported by the National Natural Science Foundation of China (grant 72604378), Kunming Medical University Clinical Research Projects (grant MR-KYLY-6), Yunnan Fundamental Research Projects (grant 202501AU070024), Yunnan Provincial Department of Education Science Research Fund Project (grant 2025J0166), Yunnan Provincial Clinical Medical Center for Blood Disease and Thrombosis Prevention and Treatment (grant 2024YNLCYXZX0253), Project of Yunnan Province Clinical Research Center for Hematologic Disease (grant 202505AJ310003), Yunnan Province Major Difficult Diseases Chinese and Western Clinical Cooperation Pilot Project—Leukemia, and the Major Collaborative Project on Clinical Treatment of Critical Illnesses by Traditional Chinese and Western Medicine.
Data Availability
Data sharing is not applicable to this article as no datasets were generated or analyzed during this study.
Authors' Contributions
Conceptualization: LY, LS, XY, ZL
Critical revision of the manuscript: ZL, YW
Funding acquisition: LY, YW
Literature review: XY, RZ
Project administration: YW
Study design: LY
Supervision: YW
Visualization: LS
Writing—original draft: LY
Writing—review and editing: LS, XY, RZ, SF
All authors contributed to the intellectual development and final approval of the manuscript.
Conflicts of Interest
None declared.
References
- Bray F, Laversanne M, Sung H, Ferlay J, Siegel RL, Soerjomataram I, et al. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA Cancer J Clin. 2024;74(3):229-263. [FREE Full text] [CrossRef] [Medline]
- Hanna TP, King WD, Thibodeau S, Jalink M, Paulin GA, Harvey-Jones E, et al. Mortality due to cancer treatment delay: systematic review and meta-analysis. BMJ. Nov 04, 2020;371:m4087. [FREE Full text] [CrossRef] [Medline]
- Dagogo-Jack I, Shaw AT. Tumour heterogeneity and resistance to cancer therapies. Nat Rev Clin Oncol. Feb 2018;15(2):81-94. [CrossRef] [Medline]
- Lotter W, Hassett MJ, Schultz N, Kehl KL, Van Allen EM, Cerami E. Artificial intelligence in oncology: current landscape, challenges, and future directions. Cancer Discov. May 01, 2024;14(5):711-726. [CrossRef] [Medline]
- Topol EJ. High-performance medicine: the convergence of human and artificial intelligence. Nat Med. Jan 2019;25(1):44-56. [CrossRef] [Medline]
- Bera K, Schalper KA, Rimm DL, Velcheti V, Madabhushi A. Artificial intelligence in digital pathology - new tools for diagnosis and precision oncology. Nat Rev Clin Oncol. Nov 2019;16(11):703-715. [FREE Full text] [CrossRef] [Medline]
- Lipkova J, Chen RJ, Chen B, Lu MY, Barbieri M, Shao D, et al. Artificial intelligence for multimodal data integration in oncology. Cancer Cell. Oct 10, 2022;40(10):1095-1110. [FREE Full text] [CrossRef] [Medline]
- Bhinder B, Gilvary C, Madhukar NS, Elemento O. Artificial intelligence in cancer research and precision medicine. Cancer Discov. Apr 2021;11(4):900-915. [FREE Full text] [CrossRef] [Medline]
- Shah NH, Entwistle D, Pfeffer MA. Creation and adoption of large language models in medicine. JAMA. Sep 05, 2023;330(9):866-869. [CrossRef] [Medline]
- Huang J, Xu Y, Wang Q, Wang QC, Liang X, Wang F, et al. Foundation models and intelligent decision-making: progress, challenges, and perspectives. Innovation (Camb). May 12, 2025;6(6):100948. [FREE Full text] [CrossRef] [Medline]
- Ferber D, El Nahhas OS, Wölflein G, Wiest IC, Clusmann J, Leßmann ME, et al. Development and validation of an autonomous artificial intelligence agent for clinical decision-making in oncology. Nat Cancer. Aug 2025;6(8):1337-1349. [CrossRef] [Medline]
- Lekadir K, Frangi AF, Porras AR, Glocker B, Cintas C, Langlotz CP, et al. FUTURE-AI: international consensus guideline for trustworthy and deployable artificial intelligence in healthcare. BMJ. Feb 05, 2025;388:e081554. [FREE Full text] [CrossRef] [Medline]
- Chang JS, Kim H, Baek ES, Choi JE, Lim JS, Kim JS, et al. Continuous multimodal data supply chain and expandable clinical decision support for oncology. NPJ Digit Med. Feb 27, 2025;8(1):128. [FREE Full text] [CrossRef] [Medline]
- Rehburg F, Graefe A, Hübner M, Thun S. How interoperability can enable artificial intelligence in clinical applications. Stud Health Technol Inform. Aug 22, 2024;316:596-600. [CrossRef] [Medline]
- Esteva A, Robicquet A, Ramsundar B, Kuleshov V, DePristo M, Chou K, et al. A guide to deep learning in healthcare. Nat Med. Jan 2019;25(1):24-29. [CrossRef] [Medline]
- Manz CR, Schriver E, Ferrell WJ, Williamson J, Wakim J, Khan N, et al. Association of remote patient-reported outcomes and step counts with hospitalization or death among patients with advanced cancer undergoing chemotherapy: secondary analysis of the PROStep randomized trial. J Med Internet Res. May 17, 2024;26:e51059. [FREE Full text] [CrossRef] [Medline]
- Chen RJ, Ding T, Lu MY, Williamson DF, Jaume G, Song AH, et al. Towards a general-purpose foundation model for computational pathology. Nat Med. Mar 2024;30(3):850-862. [CrossRef] [Medline]
- Xu H, Usuyama N, Bagga J, Zhang S, Rao R, Naumann T, et al. A whole-slide foundation model for digital pathology from real-world data. Nature. Jun 2024;630(8015):181-188. [FREE Full text] [CrossRef] [Medline]
- Azizi S, Culp L, Freyberg J, Mustafa B, Baur S, Kornblith S, et al. Robust and data-efficient generalization of self-supervised machine learning for diagnostic imaging. Nat Biomed Eng. Jun 2023;7(6):756-779. [CrossRef] [Medline]
- Selvaraju RR, Cogswell M, Das A, Vedantam R, Parikh D, Batra D. Grad-CAM: visual explanations from deep networks via gradient-based localization. Int J Comput Vis. Oct 11, 2020;128(2):336-359. [CrossRef]
- Omidi J. Integrative machine learning-driven prioritization of ceRNA networks in adrenocortical carcinoma. Intell Oncol. Apr 2026;2(2):100056. [CrossRef]
- Keyl J, Keyl P, Montavon G, Hosch R, Brehmer A, Mochmann L, et al. Decoding pan-cancer treatment outcomes using multimodal real-world data and explainable artificial intelligence. Nat Cancer. Feb 2025;6(2):307-322. [CrossRef] [Medline]
- Ismayilov R, Aktas A, Gencoglu EA, Oguz A, Altundag O, Akcali Z. Staging prostate cancer with AI: a comparative study of large language models and expert interpretation on PSMA PET-CT reports. Mol Imaging Biol. Feb 2026;28(1):93-105. [CrossRef] [Medline]
- Chen D, Huang RS, Jomy J, Wong P, Yan M, Croke J, et al. Performance of multimodal artificial intelligence chatbots evaluated on clinical oncology cases. JAMA Netw Open. Oct 01, 2024;7(10):e2437711. [FREE Full text] [CrossRef] [Medline]
- Prinster D, Mahmood A, Saria S, Jeudy J, Lin CT, Yi PH, et al. Care to explain? AI explanation types differentially impact chest radiograph diagnostic performance and physician trust in AI. Radiology. Nov 2024;313(2):e233261. [CrossRef] [Medline]
- Omar M, Sorin V, Collins JD, Reich D, Freeman R, Gavin N, et al. Multi-model assurance analysis showing large language models are highly vulnerable to adversarial hallucination attacks during clinical decision support. Commun Med (Lond). Aug 02, 2025;5(1):330. [FREE Full text] [CrossRef] [Medline]
- Geukes Foppen RJ, Morkūnas M, Traverso A, Dienstmann R. AI assistance in tumor multidisciplinary teams. ESMO Real World Data Digit Oncol. Feb 12, 2026;11:100684. [FREE Full text] [CrossRef] [Medline]
- Carl N, Hetz MJ, Wies C, Haggenmüller S, Winterstein JT, Mangold MH, et al. Enhancing clinicians' trust in large language models via transparent source attribution: a randomized controlled evaluation in uro-oncology. Eur J Cancer. Jan 17, 2026;233:116168. [FREE Full text] [CrossRef] [Medline]
- Gerke S, Simon DA, Roman BR. Liability risks of ambient clinical workflows with artificial intelligence for clinicians, hospitals, and manufacturers. JCO Oncol Pract. Mar 2026;22(3):357-361. [CrossRef] [Medline]
- Hamamoto R, Koyama T, Takahashi S, Yasuda T, Kobayashi K, Akagi Y, et al. Implementing generative artificial intelligence in precision oncology: safety, governance, and significance. J Hematol Oncol. Feb 09, 2026;19(1):14. [FREE Full text] [CrossRef] [Medline]
Abbreviations
| AJCC: American Joint Committee on Cancer |
| CHAARTED: Chemohormonal Therapy Versus Androgen Ablation Randomized Trial for Extensive Disease |
| DICOM: Digital Imaging and Communications in Medicine |
| EMR: electronic medical record |
| ETL: extract-transform-load |
| GLOBOCAN: Global Cancer Observatory |
| GTEx: Genotype-Tissue Expression |
| HIS: hospital information system |
| HL7 FHIR: Health Level Seven Fast Healthcare Interoperability Resources |
| LIS: laboratory information system |
| LLM: large language model |
| MDT: multidisciplinary team |
| NLP: natural language processing |
| PACS: picture archiving and communication system |
| PSMA PET-CT: prostate-specific membrane antigen positron emission tomography–computed tomography |
| RAG: retrieval-augmented generation |
| TCGA-ACC: The Cancer Genome Atlas adrenocortical carcinoma |
Edited by M Balcarras; submitted 04.Jun.2026; peer-reviewed by R Ismayilov, J Omidi; comments to author 13.Aug.2026; revised version received 19.Aug.2026; accepted 20.Aug.2026; published 15.Sep.2026.
Copyright©Liuyang Yang, Liyu Shan, Xiangmei Yao, Renbin Zhao, Zengzheng Li, Shuai Feng, Yajie Wang. Originally published in JMIR Cancer (https://cancer.jmir.org), 15.Sep.2026.
This is an open-access article distributed under the terms of the Creative Commons Attribution License (https://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work, first published in JMIR Cancer, is properly cited. The complete bibliographic information, a link to the original publication on https://cancer.jmir.org/, as well as this copyright and license information must be included.

